Papers with author profiling
SNAP-BATNET: Cascading Author Profiling and Social Network Graphs for Suicide Ideation Detection on Social Media (N19-3)
Copied to clipboard
Rohan Mishra, Pradyumn Prakhar Sinha, Ramit Sawhney, Debanjan Mahata, Puneet Mathur, Rajiv Ratn Shah
| Challenge: | Suicide is a leading cause of death among youth worldwide and currently only uses text-based cues to detect suicidal ideation. |
| Approach: | They propose a deep learning based model to extract text-based features from tweets and a novel Feature Stacking approach to combine other community-based information. |
| Outcome: | The proposed model outperforms existing models on an annotated dataset of tweets using a three-phase strategy and proposes a novel Feature Stacking approach to combine other community-based information such as historical author profiling and graph embeddings. |
Reusable workflows for gender prediction (L18-1)
Copied to clipboard
| Challenge: | Existing systems for author profiling (AP) modeling require extensive feature engineering and testing. |
| Approach: | They propose to implement a system for author profiling (AP) modeling that reduces the complexity and time of building a sophisticated model for a number of different AP tasks. |
| Outcome: | The proposed model achieves comparable results to state of the art models for cross-genre gender prediction, but lags when genre of test set is different from genre of train set. |
SpanEmo: Casting Multi-label Emotion Classification as Span-prediction (2021.eacl-main)
Copied to clipboard
| Challenge: | Current approaches to ER ignore potential ambiguities, in which multiple emotions overlap. |
| Approach: | They propose a model "SpanEmo" which casts multi-label emotion classification as span-prediction and introduces a loss function focused on modelling multiple co-existing emotions in a sentence. |
| Outcome: | The proposed model can predict multiple co-existing emotions in a sentence and improve model performance and learning meaningful associations between labels and words in the sentence. |
An Algerian Corpus and an Annotation Platform for Opinion and Emotion Analysis (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, there are more than 4 billion Internet users worldwide . the number of social media users in Algeria has tripled over a year . |
| Approach: | They propose a platform for crowdsourcing annotation of tweets at different levels of granularity. |
| Outcome: | The proposed platform can be used to create the largest Algerian dialect subjectivity lexicon of about 9,000 entries. |
Enhancing Sentence Embedding with Generalized Pooling (C18-1)
Copied to clipboard
| Challenge: | Existing methods for learning sentence embedding are limited, but still need to be improved. |
| Approach: | They propose a vector-based multi-head attention model that uses special cases of max pooling, mean pooling and scalar self-attention. |
| Outcome: | The proposed model improves on natural language inference, author profiling, and sentiment classification tasks. |
Author Profiling from Facebook Corpora (L18-1)
Copied to clipboard
| Challenge: | Existing studies on author profiling focus on age and gender, and use only English text. |
| Approach: | They propose to model author profiling from a Brazilian Portuguese corpus using standard gender and age prediction tasks and two less-known alternatives: predicting an author's degree of religiosity and IT background status. |
| Outcome: | The proposed tasks are based on a Brazilian Portuguese corpus and are compared with other languages and tasks. |
UMUTextStats: A linguistic feature extraction tool for Spanish (2022.lrec-1)
Copied to clipboard
| Challenge: | Feature Engineering is the application of domain knowledge to build efficient machine learning models. |
| Approach: | a team of researchers has developed a linguistic extraction tool for Spanish . the tool uses linguistic features and embeddings to build efficient machine learning models . |
| Outcome: | UMUTextStats is a linguistic extraction tool for Spanish . it has been validated in infodemiology, hate-speech detection, author profiling, authorship verification, humour or irony detection, among others. |
Preventing Author Profiling through Zero-Shot Multilingual Back-Translation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Documents as short as a single sentence may reveal sensitive information about authors . style transfer is effective but a number of current methods cause a drop in down-stream utility . |
| Approach: | They propose a method to remove sensitive information from documents by multilingual back-translation using off-the-shelf translation models. |
| Outcome: | The proposed method lowers adversarial gender and race prediction by 22% while retaining 95% of original utility on downstream tasks. |
Profiling-UD: a Tool for Linguistic Profiling of Texts (2020.lrec-1)
Copied to clipboard
| Challenge: | Profiling–UD is a text analysis tool that can be used to characterize language variation from different perspectives. |
| Approach: | They introduce Profiling–UD, a text analysis tool inspired to the principles of linguistic profiling that can support language variation research from different perspectives. |
| Outcome: | The proposed tool is specifically designed to be multilingual since it is based on the Universal Dependencies framework. |